Decomposes AI workloads into smaller tasks routed to optimal models and chips, continuously optimizing inferen