
Move your team off shared and per-user virtual API keys and onto identity-based access for your AI gateway - no keys to generate, distribute, or rotate for human users.

A general-purpose guide to putting any self-hosted model server - vLLM, Ollama, llama.cpp, LM Studio, LocalAI, or a private downstream gateway - behind Pangolin's identity-aware AI gateway.
Meet a Pangolin engineer to talk zero trust remote access, open-source, and better access control for your environment.

Access your NVIDIA DGX Spark from any laptop without SSH, a VPN, or an open port. Run Ollama on the Spark and reach it through Pangolin's identity-aware AI gateway, from the same endpoint you use for cloud models.

SSH to private servers through Pangolin with a browser terminal or private CLI over a scoped tunnel. Automatic user provisioning via PAM, without opening port 22 or distributing static keys.
© 2026 Fossorial Inc.
All systems operational