
Together AI
Products · 5 tracked products
Batch Inference
Inference service designed for large-scale asynchronous batch processing workloads. Scales single models to handle massive token data volumes efficiently, optimizing throughput for non-real-time tasks.
Dedicated Model Inference
Model inference service with exclusive hardware resources. Provides reserved throughput and a 99% availability SLA while maintaining compatibility with existing APIs for seamless integration.
Fine-Tuning Service
Open-source model fine-tuning service leveraging cutting-edge research techniques. Improves accuracy, reduces hallucinations, and controls model behavior to meet specific application requirements and safety standards.
GPU Clusters
Self-service compute service offering instant clusters up to thousands of GPUs. Deeply optimized via Together Kernel Collection to deliver high-performance training and inference capabilities efficiently.
Serverless Inference
Open-source model hosting inference service with pay-per-call pricing. Enables production-grade large model execution without infrastructure management, simplifying deployment workflows for developers and enterprises alike.
Physix Frontier