Managed inference platform for deploying and autoscaling models behind APIs.
Are you the maintainer?
Claim this page →Baseten
Managed inference platform for deploying and autoscaling models behind APIs.
01 / About
What Baseten is.
02 / Discussion CREDIBILITY-GATED
Discussion
Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.
- No discussions yet.
03 / Related
More around Baseten.
Similar tools
PlatformModal
DeploymentInference
modal-labs
Serverless GPU/CPU cloud for inference, sandboxes and agent workloads.
serverlessgpusandboxautoscaling
RunPod
DeploymentTrainingInference
GPU cloud offering Pods, autoscaling Serverless workers, multi-node Clusters and a repo-backed Hub for one-click model endpoints.
gpucloudserverlessautoscaling
ToolLLM D
DeploymentCoding
llm-d
A high-performance distributed inference serving stack optimized for production deployments on Kubernetes
kubernetesvllmkv-cachegpu
ToolLMCache
DeploymentCoding
lmcache
A KV Cache Management Layer for Scalable LLM Inference.
kv-cachevllmgpupytorch
ToolTruss
DeploymentInference
basetenlabs
Open-source format and CLI for packaging a model as a production API.
gpudockerclivllm
Fal
InferenceGenerative MediaDeployment
fal-ai
Fast generative-media inference (image, audio, video) via fal serverless models.
serverlessgpu
04 / Build
Build with Baseten.
Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.
Get Baseten →