Pricing
Konduktor
GPU cluster scheduling & management
- Priority scheduling with preemption
- RDMA-native multi-node training
- Multi-cloud job queuing
- GPU & networking health checks
- Observability: Prometheus, Grafana, logs
- Deploys in your VPC: AWS, GCP, Azure, bare-metal, and more...
Pluto
Experiment tracking
- Free tier: 2 seats, 2 GB storage
- Open source, self-host in about a minute
- W&B and Neptune-compatible API
- Sub-second UI at 3,000+ runs
- Dual-logging for safe migration
- Free for academics with a .edu email
- Enterprise tier for more than 10 seats
Why teams run Konduktor instead of rolling their own
Hardware fails. Runs survive.
Health checks catch bad GPUs early, cordon the node, and restart your job automatically. Built from managing 5M+ GPU hours across nearly every cloud.
Every layer instrumented
Telemetry across the whole cluster in Prometheus and Grafana: GPUs, nodes, and the RDMA fabric we set up for you.
You don't maintain Kubernetes
Konduktor owns the Kubernetes, RBAC, and API server. Frontend tools leave that operational burden to you.
One control plane across clouds
Queue and burst across clusters in any cloud from a single interface.
Building at a YC company? Get Trainy free or discounted.
YC companies get $500K+ in free cloud credits, and we help you turn them into trained, deployed models. Current-batch perks:
- 1 free performance-tuning consultation
- 50% off Konduktor for 6 months, or free if you have under 5 nodes
- 1 year free of Pluto's hosted Pro tier
Scaled to hundreds of GPUs on demand, reducing infrastructure costs by 50%.
Frequently Asked Questions
Konduktor deploys into your own cloud (BYOC) with a platform fee plus per-node pricing that scales sub-linearly as your cluster grows. Every deployment is sized to your cluster and support needs, so we quote it on a quick call. Book a demo and we'll walk you through it.
Pluto has a free hosted tier with 2 seats and 2 GB of storage. The hosted Pro tier is $250 per seat/month and includes up to 10 seats, 10B data points/month, 10 TB storage, unlimited tracked hours, and run forking. Self-hosting is free: Pluto is open source with full source access.
Yes. Pluto is open source. Self-host on your own servers with full source access, and no vendor can ever sunset your experiment tracking.
Yes. If you're building at a YC company, you can get Trainy free or discounted. See the For startups section above for current-batch perks.
No. For most of our customers, we help them pick a cloud provider offering that makes the most sense for their specific use case. We then assist with hardware validation to ensure they are getting the promised performance. If you already have a reserved GPU cluster, our solution can be deployed in the cloud or on-prem. For startups, we can help you go from cloud credits to a functional multi-node training setup with high bandwidth networking in < 20 mins.
Trainy offers all of the resource sharing and scheduling benefits of Slurm with much more. Teams get better workload isolation via containerization, integrated observability, and improved robustness with comprehensive health monitoring.
The first step to reducing GPU spend is cutting idle time. If you have a reserved cluster, this means having a fault-tolerant scheduler in place. A scheduler allows your team to maintain a workload queue and keep your GPUs busy 24/7, while fault-tolerance ensures that GPU failures do not require manual restarts. New and restarted workloads are placed on healthy nodes, even if they fail in the middle of the night. Once idle time has been minimized, step 2 is to look at your workload efficiency.
Yes. Konduktor gives you a single submission interface across clusters in different clouds and regions, so your team can queue and route jobs wherever capacity is available. Each individual job runs on one cluster, but you manage and burst across all of them from one place.
Kubernetes gives AI teams higher ROI on the same pool of compute. All top-tier AI research teams (OpenAI, Meta, etc.) have similar systems in place. With automated scheduling and cleanup of queued workloads, AI engineers never have to worry about GPU availability or compatibility. On the other hand, decision makers get improved visibility and control into their team's cluster usage and can make informed purchasing decisions.
Submitting jobs in Trainy's platform is done via a simple yaml file that can work across clouds. You just need to enter your existing torchrun or equivalent launch command and our platform handles the rest. Read our docs for more details.
Most Trainy customers stream data into their GPU cluster from an object store such as Cloudflare R2. In the longer term, we are looking at distributed file system integrations, but this does not exist today.
The earlier, the better. When your company is exploring gen AI applications, we help you run large-scale experiments cost-effectively. When the time comes to choose a cloud provider, we work with you to navigate cloud provider offerings, and ensure you are getting maximum performance.