
Virtual API keys solve model routing and budgeting well, but they were never built to be an identity system. Here's where that gap shows up in self-hosted LiteLLM and Bifrost deployments, and what to do instead.

A general-purpose guide to putting any self-hosted model server - vLLM, Ollama, llama.cpp, LM Studio, LocalAI, or a private downstream gateway - behind Pangolin's identity-aware AI gateway.
Meet a Pangolin engineer to talk zero trust remote access, open-source, and better access control for your environment.

Run open-weight models like DeepSeek, Qwen, and Kimi K2 on your own hardware, then reach them securely from anywhere through Pangolin's AI gateway - no API keys, no data leaving your infrastructure.

Give engineers and coding agents access to a self-hosted vLLM or Ollama cluster from one endpoint, authenticated by identity instead of API keys passed around a team.

Access your NVIDIA DGX Spark from any laptop without SSH, a VPN, or an open port. Run Ollama on the Spark and reach it through Pangolin's identity-aware AI gateway, from the same endpoint you use for cloud models.

Pangolin remote nodes let you keep the traffic edge on infrastructure you control while Pangolin Cloud handles DNS, certificates, health checks, and failover.
Get a detailed look at how Pangolin provides secure remote access for edge infrastructure, device fleets, and industrial environments.

Learn about our new remote node self-hosted offering, which combines the best of self-hosted and cloud solutions.
© 2026 Fossorial Inc.
All systems operational