---
title: "Prism — Adaptive AI Gateway"
description: "Prism is a self-hosted adaptive AI gateway for agents and teams."
url: https://prism.tm/
---

PRISM — ADAPTIVE AI GATEWAY

# One gateway for your models.

Connect supported AI providers once. Prism gives your applications one API, selects a suitable route, switches around unavailable or exhausted accounts, and keeps cost, access and diagnostics in one place.

[Get started](https://prism.tm/#editions) Read the docs

One integration

Applications and agents use one API instead of separate provider integrations.

Adaptive routing

Prism selects among eligible models, providers and accounts using policy and live signals, then performs bounded failover.

Operational control

Operators manage credentials, quotas, budgets, health, cost and diagnostics centrally.

Self-hosted

You run the gateway and control its configuration and operational data.

01 / 05 | ONE INTEGRATION

## One API instead of separate provider integrations.

Point your applications and agents at your Prism address and use the request shapes they already speak. Prism resolves the route behind it.

REQUEST

```
curl https://[your-prism-host]/v1/chat/completions \
  -H "Authorization: Bearer $PRISM_KEY" \
  -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
```

● Prism picks the model, provider and account

Endpoints

/v1/chat/completions  
/v1/responses  
/v1/models

Let Prism choose

Name a model, or leave the choice to a selector:

auto fast smart cheap local

For agents and teams

Agents keep their own work loop — goals, tools and workspace. They call Prism as their AI gateway.

02 / 05 | ADAPTIVE ROUTING

## The right available route for each request.

Prism chooses the right available model, provider and account for each request according to your policy and live operating signals.

-   Policy first your rules decide which routes are eligible
-   Live signals capability, health, quota and cost
-   Bounded failover switches around unavailable or exhausted accounts

ROUTE

request model: auto

\[provider A\] · \[account 1\] quota exhausted

● \[provider A\] · \[account 2\] served

Illustration of an account switch

03 / 05 | OPERATIONAL CONTROL

## Cost, access and diagnostics in one place.

Manage credentials, quotas, budgets, health, cost and diagnostics centrally. Prism can select lower-cost eligible routes and help reduce avoidable external spend.

SPEND ESTIMATE LIVE

ALL TIME

$\[0000.00\]

30 DAYS

$\[0000.00\]

7 DAYS

$\[000.00\]

answered no answer illustration

04 / 05 | SELF-HOSTED

## Your gateway, on your machine.

You run the gateway and control its configuration and operational data. Requests you send on to a provider are processed under that provider's own terms.

1.  01
    
    Install
    
    Run Prism on your own Linux or macOS machine.
    
2.  02
    
    Connect providers
    
    Connect supported AI providers and accounts once.
    
3.  03
    
    Send requests
    
    Point your applications and agents at one API.
    

USE CASES

## What people run through Prism.

Coding agents on one endpoint

Point coding agents, scripts and IDE assistants that accept a custom OpenAI-compatible endpoint at Prism, instead of configuring every provider in every tool.

one API · selectors · keys per agent

Keep working when a provider fails

When a provider is down, rate-limited or out of funds, Prism moves the request to another eligible route and tells you which one served it.

failover · funding state · attempt headers

Use the subscriptions you already have

Connect subscription accounts through OAuth sign-in and route your agents through them alongside API keys.

OAuth accounts · account switching

Mix local and cloud models

Run routine work on local models when they are eligible and keep cloud models for the rest. Mark a request private to forbid cloud calls for it.

local runtimes · local selector · private constraint

Keep spend under control

Give each agent its own key, see cost and tokens for every request, and set budgets that warn, downgrade, switch to local or reject.

budgets · cost per request · request log

A model runtime for your application

Embed Prism as the model layer of your own product: your application calls one API and Prism handles providers, failover and cost.

one API · routing · operational controls

Teams on one governed gateway COMING SOON

Share approved provider resources with a team, apply organizational policy and budgets, and keep an audit record.

Prism Enterprise

WHY PRISM

## Why Prism.

API keys and subscriptions, one API

Connect provider API keys and subscription accounts, and route through both from one endpoint.

One binary, nothing else to run

A single Go binary. No Python, no Docker, no database server to operate.

No silent substitution

Every attempt is disclosed in response headers, so you always know which provider served a request.

Knows when a provider runs dry

Providers that run out of funds are set aside and probed again automatically.

Free for personal use

Prism Personal has no limit on your agents, and Prism takes no share of your provider spend.

Local and cloud under one policy

Local models when they are eligible, cloud when needed, and a private flag that forbids cloud calls.

[See how Prism compares](https://prism.tm/compare/)

FEATURES

## What the gateway does.

API

-   OpenAI-compatible endpoints: chat completions, Responses, embeddings, rerank and models
-   Streaming and non-streaming responses
-   Pick a concrete route as `provider::model`, or let a selector choose: `auto`, `fast`, `smart`, `cheap`, `local`
-   Chat Completions requests can be served through an upstream Responses endpoint

Routing

-   Quality floor first, then cost: among models above the floor, Prism ranks by quality per cost; `cheap` takes the lowest-cost eligible model
-   Failover across providers on rate limits, billing errors, timeouts and outages
-   No silent substitution: every attempt is disclosed in response headers
-   Providers that run out of funds are set aside and probed again every 15 minutes by default
-   A model denylist for automatic selection
-   Local models when eligible; a `private` constraint forbids cloud calls

Providers and accounts

-   Add a provider with just an API key; models are discovered automatically
-   Presets for OpenAI, Anthropic, Google Gemini, OpenRouter, Groq, DeepSeek, Together AI, Fireworks AI, DeepInfra and Token Broker
-   Connect subscription accounts through OAuth sign-in; tokens are stored encrypted and refreshed automatically
-   Local runtimes: Ollama, llama-server, vLLM, TEI
-   An egress proxy per provider, HTTP or SOCKS5

Control and cost

-   Budgets with actions: warn, downgrade, local only, reject
-   Cost and tokens recorded for every request
-   Request log and usage history
-   Keys for your agents: issue, rotate, revoke
-   Status endpoint and Prometheus metrics
-   Web console

Deployment

-   A single Go binary: no Python, no Docker, no database server
-   Linux with systemd and macOS with launchd
-   Listens on localhost by default
-   One-command installer

Optional, off by default

-   Local-first orchestration: result checks and at most one escalation to a stronger model
-   A verifier in measurement-only mode
-   Fleet telemetry, opt-in

05 / 05 | UNDER LOAD

## A router you cannot see in the latency.

Prism sits between your agents and the model providers and decides where each request goes. We measured what that costs: how much it adds to a response, how many requests and open streams an ordinary 4 vCPU machine carries, and how much memory it takes.

0.4 ms

Prism's own latency per request. Providers answer in seconds; this disappears in them.

2,000

concurrent streaming responses on 4 vCPU. Every one completed; first token in 49 ms.

~260 MB

of memory under sustained load. It does not grow with the number of requests: half a million in a row added nothing.

4,500/s

requests per second on 4 vCPU, each one accounted for: cost, tokens, request log. Not a single record lost.

Measured on 27–28 September 2026 on 4 logical cores of an Intel Core i5-13500, with one provider and one model; the provider answers instantly so that the numbers show only Prism's cost. The figures describe that machine; other hardware and disks will give other numbers.

How we measured

EDITIONS

## Personal or Enterprise.

PRISM PERSONAL

Free

Run one AI gateway for your own agents and applications. Connect your providers and accounts, route through one API, and manage personal usage and cost without a Prism subscription.

-   Up to 3 people, including the owner
-   No limit on your agents, application credentials or provider accounts
-   Routing, providers, model catalog, diagnostics, personal cost and budgets

[Get started](https://prism.tm/#editions)

PRISM ENTERPRISE COMING SOON

$50 per team / month · up to 8 users

Give teams one governed gateway to AI providers and accounts. Apply access policy and budgets centrally, share approved resources, and retain the audit record required to operate AI across an organization.

-   Teams and managed membership
-   Organizational access policy and shared resource administration
-   Team budgets and corporate audit

Contact us

Free refers to Prism itself. Provider usage and infrastructure may still cost money.

## Smarter choice.

Prism is a self-hosted adaptive AI gateway for agents and teams.

[Get started](https://prism.tm/#editions) Read the docs
