
Together AI
Product Matrix

Batch Inference
Inference service designed for large-scale asynchronous batch processing workloads. Scales single models to ha

Dedicated Model Inference
Model inference service with exclusive hardware resources. Provides reserved throughput and a 99% availability

Fine-Tuning Service
Open-source model fine-tuning service leveraging cutting-edge research techniques. Improves accuracy, reduces

GPU Clusters
Self-service compute service offering instant clusters up to thousands of GPUs. Deeply optimized via Together

Serverless Inference
Open-source model hosting inference service with pay-per-call pricing. Enables production-grade large model ex
Physix Frontier