Reach private gateways over the Pangolin client, reach self-hosted models over site connectors, and keep identity, budgets, and logs on every call.
Keyless access over the Pangolin client tunnel
Private AI Gateway resources stay off the public internet. Connected clients already prove identity, so Claude Code, Codex, OpenCode, and Gemini CLI call the gateway without minting another provider secret.
Tunnel to self-hosted models behind your firewall
Reach self-hosted models and other private endpoints over site tunnels, then serve them alongside OpenAI, Anthropic, Gemini, and other cloud providers on one URL.
Budgets, session logs, and usage analytics
Cap spend or tokens by provider, model, resource, role, or key. Store transcripts and chart cost, tokens, and request volume across every client and upstream you connect.
Built like the rest of Pangolin.
AI Gateway is a protocol-aware resource on the same platform as your apps, databases, SSH, and other resources, not a separate LLM proxy to operate beside it.
Manage AI like other infrastructure
Model access belongs next to the databases, SSH, RDP, and internal apps you already protect in Pangolin. Treat the gateway like any other resource, with the same users, roles, and domains.
Grant users and roles instead of managing API keys
Skip API keys entirely with the Pangolin client, or use identity keys to identify the caller. In both cases, users must log in with your existing identity provider first.
Speak each provider's native API format
Clients keep speaking the protocol they already expect. Pangolin translates and routes each request to the right upstream, so you point tools at one hostname instead of wiring every format by hand.
Run multiple gateway endpoints with their own URL
Traditional LLM gateways expose one URL for the whole instance. In Pangolin, each AI Gateway is its own resource, so you can run many hostnames with different providers, models, and access rules.
Tunnel AI the same way you tunnel everything else.