This changes capacity, network requirements and failure domains.
MODEL SERVINGENTERPRISE AI ARCHITECTURE
Model Parallelism
Large models may span multiple GPUs using tensor, pipeline or other parallel techniques.
A four-GPU tensor-parallel replica requires all four GPUs healthy before it can serve.
How does one hardware failure affect the replica and pool?
REMEMBERParallelism increases both capability and coordination dependency.
No uploads · No company data · No account required