LLM 网关与推理服务
共 28 款LLM Gateways & Serving类AI工具
分类导读
LLM 网关和推理服务连接应用与一个或多个模型端点,用于统一调用、路由、限流、观测和部署。选择时应基于真实请求负载验证延迟、稳定性与成本。
选择维度
- 模型提供方、开源模型、API 协议、流式响应和工具调用兼容性
- 路由、缓存、批处理、限流、重试、回退和并发控制能力
- 密钥隔离、数据保留、私有部署、访问策略和审计日志
- 吞吐、首字延迟、可用性、GPU 或请求计费和迁移成本
使用提醒
上线前应使用目标模型、上下文长度和并发模式进行压测;为超时、限流、上游故障和模型切换设置可观察的回退策略,避免静默改变结果。
来源与更新
页面按本站 2026-08-27 的 28 个服务入口整理;协议、支持模型、数据策略、区域和价格以各平台官方文档为准。
相邻分类
- fal.ai Freemium ★4.93Fastest generative AI platform for developers — 1,000+ image, video, audio, and 3D models with optimized real-time inference. Default home for FLUX, SAM, MuseTalk.
- FluidStack Paid ★4.93FluidStack: On-demand GPU servers for ML, rendering, and general compute tasks.
- Ollama Free ★4.93Ollama is a local LLM runtime that downloads, runs, and serves open models on your own hardware via a CLI and an OpenAI-compatible API.
- TOGETHER Paid ★4.93Cloud service for developers to build with open-source AI, offering APIs, distributed training systems, and leading open-source models.
- Fireworks AI Paid ★4.92High-speed, cost-efficient generative AI for product innovation with advanced fine-tuning capabilities.
- RunPod Paid ★4.91Globally distributed GPU cloud for AI tasks.
- Groq Paid ★4.87Enterprise-scale AI solutions for ultra-fast language processing and inference.
- Modal Paid ★4.86Modal offers an easy way for developers to run code in the cloud with serverless compute and containerized environments.
- Not Diamond Paid ★4.84Not Diamond is a model routing layer that selects the right LLM for each query to raise quality and cut costs.
- OpenRouter Freemium ★4.84Unified API and marketplace for the best LLMs at the best prices for any prompt.
- Sail Research Freemium ★4.84Sail Research is an inference platform that pairs low-cost model serving with stateful agent sandboxes.
- TrueFoundry Freemium ★4.84Enterprise AI gateway and platform to deploy, govern and scale LLMs, agents and MCP tools on any cloud.
- Kong Freemium ★4.83Kong is an AI connectivity platform that secures, manages, and monetizes API and AI token traffic.
- LiteLLM Freemium ★4.75Universal LLM proxy — call 100+ LLMs (OpenAI, Anthropic, Bedrock, Vertex) with one API.
- Voltage Park Paid ★4.75Voltage Park is a GPU cloud platform that rents NVIDIA H100 and Blackwell clusters on-demand or on dedicated reserve for AI training and inference.
- TensorDock Paid ★4.7Affordable and flexible GPU cloud computing for AI, ML, and rendering.
- Baseten Paid ★4.63AI-powered platform for building and deploying machine learning models.
- Vast AI Paid ★4.6AI platform for affordable and flexible GPU cloud computing.
- Kindo Paid ★4.58Kindo is the secure enterprise GenAI gateway — single SSO into multiple LLMs with policy, logging, and data-loss prevention. Drive Capital-led.
- DeepInfra Paid ★4.46DeepInfra is an inference cloud that serves open-weight AI models — Llama, DeepSeek, Qwen, Mistral — behind a pay-per-token, OpenAI-compatible API.
- FriendliAI Paid ★4.45FriendliAI is the LLM inference platform behind Friendli Container, Dedicated, and Serverless Endpoints. Competes with Together AI and Fireworks.